You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
🔴 Possibly broken: pytest==9.1.1 — pytest 9.x did not exist on PyPI at the time of writing (8.x was current). If this version doesn't exist, pip install fails hard and the workflow is red on every PR.
Suggestion: Verify the version exists; consider pinning to a known-good release, or drop the pin and use a requirements-test.txt in the repo so dependency drift is reviewed like code.
🟡 Inconsistent supply-chain pinning — setup-python is pinned to a full commit SHA, but actions/checkout@v4 floats on a tag.
Suggestion: Pin checkout to a SHA too, or (if you prefer tag-level) apply the same policy to setup-python. The inconsistency signals that the pinning convention isn't documented or enforced.
🟡 No push trigger — The check only runs on pull_request. A force-push or admin push directly to main bypasses it entirely.
Suggestion:
on:
push:
branches: [main]pull_request:
# ...
🟡 set -euo pipefail only in one of two run steps — The check_linux_repro_docs.py step has no explicit error handling. GitHub's default shell does set -e, but pipefail and nounset are not set by default.
Suggestion: Either add a reusable shell preamble, or use shell: bash + a common env setup. This matters if the script ever uses pipelines or references an unset variable.
💭 Consider cache: pip in setup-python — Even for a small test suite, it avoids re-downloading on every PR. Negligible here, but it's a one-liner that becomes valuable as the test suite grows.
📁 .github/workflows/llm-deploy.yml
🟡 full-qwen3-frontend downloads ~1.24GB on every PR — No caching. Each PR pays full bandwidth and time for the model download. Consider actions/cache on output/qwen3-full-model/, or gating this job to pull_request only for path-filtered changes and running it on workflow_dispatch/schedule otherwise.
🟡 w2-acceptance needs correct skipped-handling — if: always() + toJSON(needs) means check_w2_ci_results.py receives full-qwen3-frontend.result as "skipped" when path filters don't match. If that script treats "skipped" as "success", a PR touching only docs could pass W2 acceptance without any actual gate executing. Verify it requires "success" not just "not failure."
🟡 Mixed execution styles for unit gates — unit:frontend-ops runs pytest, unit:backend-ops runs python3 probes/.../run.py directly. The probe produces a report file but no JUnit XML. If the probe script fails silently (non-zero exit but incomplete report), the gate could mask a real failure. Ensure run.py exits non-zero on assertion failures.
💭 unit:runtime (step) has no per-step timeout — Job timeout is 60 min total; add timeout-minutes to the step too for faster feedback on hangs.
💭 Unnecessary quotes on env values — OMP_NUM_THREADS: "1" etc. GitHub Actions always stores env as strings; bare 1 is equivalent and slightly cleaner.
🟡 Suggestion — This bullet conflates two distinct policies (platform requirements + historical data preservation). Splitting them improves scannability and makes it easier to update one without touching the other:
- Qwen3 deployment (W1+) requires Linux (Ubuntu 24.04, Bash, Python 3.12,
Linux tool paths) in all public guides, PR instructions, and CI.
See [Linux reproduction contract](docs/llm-deploy-v1.0/LINUX_REPRODUCTION.md).
- Preserve the actual platform and commit identity of historical measurements;
personal dev environments are not acceptance evidence.
💭 Nit — Confirm docs/llm-deploy-v1.0/LINUX_REPRODUCTION.md exists at the referenced path before merging, or the link will be dead.
📁 benchmarks/test_regalloc/regalloc.md
🟡 Windows 指令被完全移除 — 原先的 PowerShell 命令 (Windows PowerShell) 被替换而非保留为备选。如果仓库仍有 Windows 用户,建议在注释中提一句"Windows 用户请参考 X 路径或 pwsh"。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
为 W2 建立可独立复现的完整 Qwen3 前端、基础后端和 Host 运行时验收。完整模型不再仅由小 fixture 代表:逐项审计固定 28 层 ONNX 的实际节点、IR 和权重绑定;PR 必须实际运行完整前端任务,汇总检查拒绝失败、取消、跳过和缺失结果。
本 PR 直接基于已合并主干
3bb88e8。W1 两层报告与固定 ORT 参考用例修复为独立 PR #93;两笔 PR 的修改文件不重叠,本分支已不依赖 #93 重新执行七项技术验收。共享后端/运行时的常量绑定、日志与取消清理加固随本 PR 交付。修改
Linux 交付与复现(2026-10-06)
正式交付、CI 和第二人复现统一使用 Ubuntu 24.04 / Bash / Python 3.12。Linux 是运行编译器、ORT、IR 解释器和 QEMU 的 Host;目标程序仍使用 RV64 裸机 ABI。
当前提交
f7ebbfa745ca5be31aea857425b0e97e5ce614c2;相对22bd3b58b5eed5f43bd43f334c9ba45b4c8bf3cb仅补齐通用环境安装文档,执行代码与 CI 配置未变。此次更新将 W1/W2、探测及相关 benchmark 指南统一为 Linux:scripts/run_w2_acceptance.py。执行代码提交
22bd3b5的 Linux 文档 CI 与 Topic06 已成功;CI 和 LLM Deploy 也均已成功。最新文档提交的 Linux 文档检查 已成功;其余工作流按当前 head 的 Checks 记录。本次改动不调整模型或数值算法。已有 Linux 验证与边界
下面属于历史提交
2d07ccc42936dcfde19d552bca02a324cd33821c,不能冒充本次更新的新运行:test_w2_acceptance.py在 Linux 55/55 通过、0 跳过,包含新增 SIGINT/SIGTERM 实际取消测试。LLM 确认 32/32 基础后端、28/28 两层 QEMU、3/3 显式产物集成及 20/20 次超时清理通过;两层诊断最大误差1.7285346984863281e-6 < 1e-5。完整 ONNX release 任务按 PR 条件跳过,不能算作完整 ORT 重验。数值门槛保持严格最大绝对误差与
rtol=0。完整前端解析不等于完整 0.6B IR/QEMU 前向,单步随机模型不等于可用文本生成。源码交付不包含模型、ELF、缓存或个人依赖。W1 报告可靠性修复已由 PR #93 合并。作者 Linux CI 不替代第二人的七项实际执行;团队接口确认、评审与阶段出口仍待按记录完成,关联 #92。后续完整 IR 工作由 PR #96 交付。